Papers with GRPO baseline

2 papers
Harmonizing Dense and Sparse Signals in Multi-turn RL: Dual-Horizon Credit Assignment for Industrial Sales Agents (2026.acl-industry)

Copied to clipboard

Challenge: Large language models for industrial sales require balancing long-term commercial objectives with immediate linguistic constraints such as fluency and compliance.
Approach: They propose a framework that disentangles optimization across time scales by normalizing advantages from turn-level and session-level rewards before fusion.
Outcome: The proposed framework outperforms the state-of-the-art GRPO model in conversion rate and identity detection rate.
Outcome-Grounded Advantage Reshaping for Fine-Grained Credit Assignment in Mathematical Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Group Relative Policy Optimization (GRPO) uses a coarse-grained credit assignment mechanism that propagates group-level rewards uniformly to to every token in a sequence, neglecting the varying contribution of individual reasoning steps.
Approach: They introduce Outcome-grounded Advantage Reshaping (OAR) which redistributes advantages based on how much each token influences the model’s final answer.
Outcome: Empirical results show that OAR-G outperforms GRPO on a high-fidelity attribution signal and suppresses low-impact tokens while preserving the advantage mass.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations